Papers with DPO variants

3 papers
LiPO: Listwise Preference Optimization through Learning-to-Rank (2025.naacl-long)

Copied to clipboard

Challenge: Recent work on language models with curated feedback provides promising alternatives to RLHF . multiple responses can be ranked by reward models or AI feedback, but there is no study on directly fitting upon a list of responses.
Approach: They propose a method that aligns language models with curated human feedback . they propose SLiC and DPO as promising alternatives to traditional RLHF .
Outcome: The proposed method outperforms DPO and SLiC on several preference alignment tasks with curated and real rankwise preference data.
IRPO: Implicit Policy Regularized Preference Optimization (2026.findings-eacl)

Copied to clipboard

Challenge: Recent DPOs introduce additional hyperparameters, reducing feasibility for LLM fine-tuning.
Approach: They propose an algorithm that regularizes the reward against a reference policy without extra hyperparameters to address suboptimal outcomes.
Outcome: The proposed algorithm outperforms baseline algorithms with the same hyperparameter complexity while maintaining training simplicity.
Sem-DPO: Mitigating Semantic Inconsistency in Preference Optimization for Prompt Engineering (2026.findings-acl)

Copied to clipboard

Challenge: Direct Preference Optimization (DPO) is an off-policy alternative to RL for automatic prompt engineering, but its token-level regularization leaves semantic inconsistency unchecked as prompts that win higher preference scores can still drift away from the user’s intended meaning.
Approach: They propose a variant of Direct Preference Optimization that preserves semantic consistency while maintaining its simplicity and efficiency.
Outcome: The proposed model outperforms state-of-the-art prompt optimization baselines and several DPO variants on three standard text-to-image prompt-optimization benchmarks and three language models.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations